Papers with language model distillation

3 papers
Minimal Distillation Schedule for Extreme Language Model Compression (2024.findings-eacl)

Copied to clipboard

Challenge: Existing methods for teacher assistant-based distillation require multiple trials to find the optimal teacher assistant.
Approach: They propose a method that allows scheduling of an optimal teacher assistant in just one trial . they show that student performance is positively correlated with the scale-performance tradeoff .
Outcome: The proposed method can select the optimal teacher assistant in just one trial . it can be used to compare performance of student and teacher assistants on GLUE benchmarks.
Learning to Maximize Mutual Information for Chain-of-Thought Distillation (2024.findings-acl)

Copied to clipboard

Challenge: Knowledge distillation is a technique of transferring knowledge from large, complex models to smaller ones.
Approach: They propose a method utilizing chain-of-thought distillation to transfer knowledge from large, complex models to smaller ones by maximizing mutual information of the representation features of the two tasks.
Outcome: The proposed method outperforms the state-of-the-art knowledge distillation method on four datasets.
CIF-PT: Bridging Speech and Text Representations for Spoken Language Understanding via Continuous Integrate-and-Fire Pre-Training (2023.findings-acl)

Copied to clipboard

Challenge: Speech-to-text training and language model distillation are used to bridge the representations between speech and text.
Approach: They propose a pre-training paradigm that integrates speech and text into a single frame-to-token alignment.
Outcome: The proposed paradigm outperforms the state-of-the-art model on intent classification and slot filling tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations